Papers with vision and language understanding
Borrowing Human Senses: Comment-Aware Self-Training for Social Media Multimodal Classification (2022.emnlp-main)
Copied to clipboard
| Challenge: | Social media users are using images and text to voice opinions and share ideas. |
| Approach: | They propose to use user comments to extract hinting features from user comments and explore them via self-training. |
| Outcome: | The proposed framework improves on four social media benchmarks for image-text relation classification, sarcasm detection, sentiment classification, and hate speech detection. |
M2C: Towards Automatic Multimodal Manga Complement (2023.findings-emnlp)
Copied to clipboard
| Challenge: | Multimodal manga analysis focuses on enhancing manga understanding with visual and textual features. |
| Approach: | They propose a task to enhance manga understanding with visual and textual features by providing a shared semantic space for vision and language understanding. |
| Outcome: | The proposed task provides a shared semantic space for vision and language understanding. |
Automatic Evaluation for Text-to-image Generation: Task-decomposed Framework, Distilled Training, and Meta-evaluation Benchmark (2025.acl-long)
Copied to clipboard
| Challenge: | Existing MLLMs rely on commercial models such as GPT-4o for evaluations, but they are not universally accessible. |
| Approach: | They propose a task decomposition evaluation framework based on GPT-4o to automatically construct a specialized training dataset to break down the multifaceted evaluation process into simpler sub-tasks. |
| Outcome: | The proposed framework outperforms the current state-of-the-art GPT-4o evaluation framework with over 4.6% improvement in Spearman and Kendall correlations with human judgments. |